WEBVTT
Kind: captions
Language: en

00:00:00.060 --> 00:00:03.740
Have you ever noticed that a loudspeaker is
the opposite of an eardrum?

00:00:04.180 --> 00:00:05.320
Eh, probably not.

00:00:05.320 --> 00:00:06.390
But it’s true!

00:00:06.390 --> 00:00:10.809
See our ears work by concentrating changes
in air pressure onto a small diaphragm that

00:00:10.809 --> 00:00:13.710
will move back and forth with the pressure
changes.

00:00:13.710 --> 00:00:18.560
This vibration causes stimulation in the heary
bits of the ear which your brain can, assuming

00:00:18.560 --> 00:00:22.730
you have normal hearing ability, turn into
what we perceive as sound.

00:00:22.730 --> 00:00:27.720
A loudspeaker does the opposite--its diaphragms
(the driver cones) vibrate to create pressure

00:00:27.720 --> 00:00:29.160
changes in the air.

00:00:29.160 --> 00:00:32.629
This vibration gets transferred to our eardrums
so we can hear it.

00:00:32.629 --> 00:00:36.230
We’re sticking to simple stuff today because
the rabbit hole is just too deep.

00:00:36.230 --> 00:00:39.860
All you need to know is that things vibrate,
which causes air pressure to fluctuate, which

00:00:39.860 --> 00:00:44.040
causes our eardrums to also vibrate, which
stimulates the brain so that we can perceive

00:00:44.040 --> 00:00:46.579
that vibration as sound.

00:00:46.579 --> 00:00:50.670
This channel started as a series exploring
the history of artificial sound, and it’s

00:00:50.670 --> 00:00:54.719
been over TWO YEARS since I last touched on it at all.

00:00:54.719 --> 00:00:57.640
Finally we’re finishing this up with the
introduction of

00:00:57.640 --> 00:01:00.000
DIGITAL SOUND
(emphasis added with obnoxious reverb).

00:01:00.000 --> 00:01:04.860
Since it’s been forever, let’s go over
a brief history of sound recording technologies.

00:01:04.869 --> 00:01:08.500
The first device which could reproduce a sound
recording was the phonograph.

00:01:08.500 --> 00:01:12.650
Thomas Edison’s invention consisted of an
artificial eardrum, which would vibrate along

00:01:12.650 --> 00:01:17.290
with changes in sound pressure, and with the
aid of a collecting horn, the vibration is

00:01:17.290 --> 00:01:21.190
transferred into this stylus, creating an
up-and-down motion.

00:01:21.190 --> 00:01:25.479
This carves a groove into a wax cylinder,
and the vibrating stylus creates an imprint

00:01:25.479 --> 00:01:26.910
of the sound wave.

00:01:26.910 --> 00:01:32.049
The depth of that groove becomes a literal
analog of the original sound vibrations.

00:01:32.049 --> 00:01:36.880
Then, when the stylus is run over the now
bumpy groove, the bumps cause the diaphragm

00:01:36.880 --> 00:01:41.460
to vibrate in the same way as it did when
it first made the bumps, and the result is

00:01:41.460 --> 00:01:44.170
that you hear the same sound as before.

00:01:44.170 --> 00:01:48.660
Or at least, a barely passable imitation of
that sound.

00:01:49.620 --> 00:01:54.740
(sad sounding violin music)

00:02:02.740 --> 00:02:06.119
Commercially produced
discs and cylinders were molded from master

00:02:06.119 --> 00:02:09.570
recordings, and wouldn’t wear down like
the original wax cylinders.

00:02:09.570 --> 00:02:12.070
They were played back using devices like this.

00:02:12.070 --> 00:02:15.830
This device is called a reproducer, and
for decades all phonographs were based on

00:02:15.830 --> 00:02:18.380
simple acoustic devices like this.

00:02:18.380 --> 00:02:22.860
For nearly a century, this is how artificial
sound recording technologies worked.

00:02:22.860 --> 00:02:28.040
Something (like this horn) would collect sound
waves, and recreate them onto a physical analog.

00:02:28.050 --> 00:02:33.700
Then, that physical analog could recreate
the original sound waves when played back.

00:02:33.700 --> 00:02:38.440
While it all started with simple acoustic
devices like this phonograph, eventually improvements

00:02:38.440 --> 00:02:39.440
were made.

00:02:39.440 --> 00:02:43.290
The development of the electronic microphone
was perhaps the most important.

00:02:43.290 --> 00:02:48.810
Now, sound waves cause a receiving diaphragm
to move a coil of wire around a magnet, and

00:02:48.810 --> 00:02:52.510
a voltage is produced in the wire as the diaphragm
moves.

00:02:52.510 --> 00:02:56.660
This time, sound waves are recreated as a
voltage coming from the microphone, and by

00:02:56.660 --> 00:03:01.250
amplifying this voltage and sending it into
a new record cutting device which moves its

00:03:01.250 --> 00:03:05.830
cutting stylus as a function of the voltage
it receives, a more accurate carving of the

00:03:05.830 --> 00:03:09.050
sound wave could be made into a disc or cylinder.

00:03:09.050 --> 00:03:13.290
This greatly improved the fidelity of the
recorded sound, even on acoustic reproduction

00:03:13.290 --> 00:03:14.840
devices like this.

00:03:14.840 --> 00:03:19.340
With the proliferation of radio--which I feel
I must explain is a sound transmission technology,

00:03:19.340 --> 00:03:20.740
not sound recording.

00:03:20.740 --> 00:03:23.160
Just so we don’t get confused too much here--

00:03:23.160 --> 00:03:25.860
the loudspeaker became a big deal.

00:03:25.860 --> 00:03:30.960
Loudspeakers are the opposite of microphones--instead
of producing a voltage as a reaction to a

00:03:30.970 --> 00:03:35.590
sound pressure wave moving its diaphragm,
a loudspeaker will move its diaphragm and

00:03:35.590 --> 00:03:39.260
create a pressure wave as a reaction to incoming
voltage.

00:03:39.260 --> 00:03:44.030
With loudspeakers all the rage, record players
could now use a phonograph cartridge, which

00:03:44.030 --> 00:03:46.460
acts like a microphone for records.

00:03:46.460 --> 00:03:51.650
The movement of the stylus as the groove vibrates
it generates a voltage which can be amplified

00:03:51.650 --> 00:03:53.380
to drive a loudspeaker.

00:03:53.380 --> 00:03:56.080
This gets very meta very quickly.

00:03:56.090 --> 00:04:00.520
An artificial ear turns sounds into voltage,
and a cutting stylus turns this voltage into

00:04:00.520 --> 00:04:01.990
a groove on a record.

00:04:01.990 --> 00:04:07.840
Then, a playback stylus playing the record
generates a voltage as the stylus vibrates.

00:04:07.840 --> 00:04:12.040
This voltage is then amplified to drive a
loudspeaker, which causes pressure changes

00:04:12.040 --> 00:04:16.390
in the air around the loudspeaker, which your
ears concentrate down to your eardrums, and

00:04:16.390 --> 00:04:21.070
now your real eardrums are vibrating in roughly
the same way that the original artificial

00:04:21.070 --> 00:04:24.330
eardrum moved in the microphone in the first
place.

00:04:24.330 --> 00:04:25.330
Yeah.

00:04:25.330 --> 00:04:30.350
In essence, the record becomes a way to recreate
the original pattern of voltage created by

00:04:30.350 --> 00:04:34.480
the microphone, so that the sound can be heard
again in a different place

00:04:34.480 --> 00:04:36.260
at a different time.

00:04:36.260 --> 00:04:38.880
Let’s cut out the middle bit because that’s
what’s most confusing.

00:04:38.880 --> 00:04:42.930
A microphone like this creates an electrical
signal of fluctuating intensity based on how

00:04:42.930 --> 00:04:44.690
its diaphragm moves.

00:04:44.690 --> 00:04:48.930
I can just amplify that signal and send it
straight into a loudspeaker, which will reproduce

00:04:48.930 --> 00:04:50.840
the sound in real time.

00:04:50.840 --> 00:04:54.400
Radio accomplishes this wirelessly, but the
sound isn’t recorded.

00:04:54.400 --> 00:04:59.300
To capture the sound coming from the microphone
to be played back later, it has to be converted

00:04:59.300 --> 00:05:01.770
into an analog of the signal.

00:05:01.770 --> 00:05:04.900
And that’s why it’s called analog recording
technology.

00:05:04.900 --> 00:05:09.820
No matter if it’s a record, a cassette tape,
an open reel tape, or even a wax cylinder,

00:05:09.820 --> 00:05:12.370
the sound information is recorded “doorectly”...

00:05:12.370 --> 00:05:13.370
Doorectly.

00:05:13.370 --> 00:05:14.370
Doorectly?

00:05:14.370 --> 00:05:18.470
The sound information is recorded directly
onto something, which can then be used to

00:05:18.470 --> 00:05:22.110
recreate a copy of the original sound information.

00:05:22.110 --> 00:05:25.720
That something is an analog of the original
sound waves.

00:05:25.720 --> 00:05:30.100
Improvements in sound technology were for
many years simply incremental.

00:05:30.110 --> 00:05:32.530
Wax cylinders became shellac discs.

00:05:32.530 --> 00:05:34.110
Shellac became vinyl.

00:05:34.110 --> 00:05:38.100
Magnetic recording wire allowed for a reusable,
electronic recording medium.

00:05:38.100 --> 00:05:42.600
This was improved into magnetic tape, allowing
for a high fidelity, versatile recording medium

00:05:42.600 --> 00:05:45.340
enabling multi-track recording and editing.

00:05:45.340 --> 00:05:49.400
And to improve on the noise of magnetic tape,
different particle formulations were developed,

00:05:49.400 --> 00:05:51.540
and noise reduction technologies matured.

00:05:51.540 --> 00:05:55.810
But we were still just taking some signal
from a microphone, then slapping it basically

00:05:55.810 --> 00:05:59.040
as is onto some sort of physical medium.

00:05:59.040 --> 00:06:01.420
And that medium was never perfect.

00:06:01.420 --> 00:06:04.490
Poorly made tape would cause signal dropouts.

00:06:04.490 --> 00:06:09.500
Discs would be plagued by dust and scratches,
and would slowly wear down with each play.

00:06:09.500 --> 00:06:14.010
Because the analog medium contained the sound
in its physical properties, it was inherently

00:06:14.010 --> 00:06:16.770
prone to wear, damage, and distortion.

00:06:16.770 --> 00:06:21.680
Which of course would wear down, damage, or
distort the sound recording itself.

00:06:21.680 --> 00:06:26.670
If only there were some way to encode the
sound, perhaps a way to store sound logically

00:06:26.670 --> 00:06:28.680
rather than analogously.

00:06:28.680 --> 00:06:32.910
Maybe if the signal weren’t the sound itself,
but instead were a set of instructions on

00:06:32.910 --> 00:06:37.600
how to recreate it, we could get lossless,
near-perfect sound reproduction.

00:06:37.600 --> 00:06:40.010
And thus, digital sound was born.

00:06:40.010 --> 00:06:44.770
The heart of uncompressed digital sound is
pulse-code modulation, or PCM.

00:06:44.770 --> 00:06:49.840
PCM’s roots can be traced back to the telegraph
days, but its invention as we know it today

00:06:49.840 --> 00:06:53.510
for sound came from British Engineer Alec
Reeves.

00:06:53.510 --> 00:06:57.450
I feel I must compliment Mr. Reeves on his
given name, it’s excellent.

00:06:57.450 --> 00:06:58.450
Very good.

00:06:58.450 --> 00:07:03.570
He first devised this digital method of transmitting
and receiving voice communication in 1937,

00:07:03.570 --> 00:07:06.890
though it required extremely complex circuitry
for the time.

00:07:06.890 --> 00:07:11.910
However, PCM transmission was used during
World War 2 as a way to encrypt extremely

00:07:11.910 --> 00:07:15.394
important voice conversations, such as those
between Winston Churchill

00:07:15.394 --> 00:07:17.620
and Franklin Delano Roosevelt.

00:07:17.620 --> 00:07:19.800
This encryption system was called SIGSALY,

00:07:19.800 --> 00:07:20.880
“SIGSALLY”?

00:07:20.880 --> 00:07:22.020
“SIGSALIE”?

00:07:22.460 --> 00:07:26.980
Or Project X, X System, Ciphony 1, or Green
Hornet.

00:07:26.990 --> 00:07:31.550
Anyway, Project Green Sally X System Hornet
1 was much more complicated than simple Pulse

00:07:31.550 --> 00:07:35.450
Code Modulation, but PCM was a large part
of its encryption.

00:07:35.450 --> 00:07:37.370
So how does PCM work?

00:07:37.370 --> 00:07:39.639
It’s actually simpler than it might seem
at first.

00:07:39.639 --> 00:07:43.980
It’s rather like a system for repeatedly
asking what the instantaneous amplitude of

00:07:43.980 --> 00:07:48.980
a signal is many thousands of times per second,
then simply writing that down.

00:07:48.980 --> 00:07:51.200
Let’s look at a simple sine wave.

00:07:51.200 --> 00:07:55.450
If this were to be encoded on a vinyl record,
the groove of the record would start out straight

00:07:55.450 --> 00:07:59.240
in the center, then move to the left as the
signal intensity reached peak, then it would

00:07:59.240 --> 00:08:02.530
start to move to the right, keep moving, keep
moving, and then it would pull back to the

00:08:02.530 --> 00:08:03.670
center.

00:08:03.670 --> 00:08:07.639
When it’s played back, the movement of the
stylus as the walls of the groove wiggle it

00:08:07.639 --> 00:08:10.640
back and forth will recreate this signal.

00:08:10.640 --> 00:08:14.800
And audio tape does the same thing, except
the intensity isn’t recorded as a physical

00:08:14.800 --> 00:08:18.080
movement, but as a degree of magnetization
on the tape.

00:08:18.080 --> 00:08:21.650
But with PCM, we aren’t even trying to recreate
the wave.

00:08:21.650 --> 00:08:24.980
Instead, we want to quantify it and play connect-the-dots.

00:08:24.980 --> 00:08:28.320
Let’s say I want to take 20 samples of this
waveform.

00:08:28.320 --> 00:08:31.140
OK, I’ll divide it up into 20 chunks.

00:08:31.140 --> 00:08:34.820
Now I just need to define the detail I can
have within each sample.

00:08:34.820 --> 00:08:37.040
Let’s put this on a scale of 0 to 15.

00:08:37.040 --> 00:08:38.940
That's 4 bits of resolution.

00:08:38.940 --> 00:08:42.890
Now, at each sampling point, we can take the
closest value.

00:08:42.890 --> 00:08:46.920
This sine wave can now be represented as the
following string of numbers.

00:08:46.920 --> 00:08:50.370
To get the sine wave back, we simply plot
those numbers on a graph.

00:08:50.370 --> 00:08:52.520
Then, connect the dots.

00:08:52.520 --> 00:08:54.500
Tada! A sine…

00:08:54.500 --> 00:08:55.080
wave?

00:08:55.500 --> 00:08:56.980
Well, a sloppy sine wave.

00:08:56.990 --> 00:08:59.360
But that’s only because we weren’t very
specific.

00:08:59.360 --> 00:09:03.770
We only took 20 samples, and each one could
only be one of 16 values.

00:09:03.770 --> 00:09:07.850
But now we know the two most crucial parts
of digital sound--the sample rate and the

00:09:07.850 --> 00:09:09.270
bit depth.

00:09:09.270 --> 00:09:13.540
Perhaps the most common sample rate and bit
depth of digital sound is 44.1 kilohertz,

00:09:13.540 --> 00:09:15.000
16 bits.

00:09:15.000 --> 00:09:21.660
This means that every second, 44,100 samples
are taken, and each sample can be one of 65,536

00:09:21.660 --> 00:09:24.720
values, or 2 to the power of 16.

00:09:24.720 --> 00:09:29.360
And that’s how devices like this, a Tascam
DR-05, record sound.

00:09:29.360 --> 00:09:34.240
It’s looking at the voltage coming from
the microphone, and taking precise measurements.

00:09:34.240 --> 00:09:40.610
Every 44.1 thousandth of a second, it takes
a voltage reading, and, well, writes it down.

00:09:40.610 --> 00:09:45.870
It’s furiously quantifying and logging the
voltage it measures with 16 bits of accuracy,

00:09:45.870 --> 00:09:49.970
and the result is a string of numbers that
logically represent the shape of the sound

00:09:49.970 --> 00:09:53.430
waves that exerted pressure on the microphone’s
diagram.

00:09:53.430 --> 00:09:54.610
Pretty neat, huh?

00:09:54.610 --> 00:09:58.470
And it can actually write down two numbers
at a time, since this has two microphones

00:09:58.470 --> 00:10:00.180
and records in stereo.

00:10:00.180 --> 00:10:04.780
Inside this recorder is what’s called an
analog-to-digital converter, or ADC.

00:10:04.780 --> 00:10:09.450
The “ADC” is the actual device responsible
for creating the stream of samples.

00:10:09.450 --> 00:10:13.220
It takes the analog signal coming from the
microphones themselves and converts it into

00:10:13.220 --> 00:10:14.930
a stream of discrete numbers.

00:10:14.930 --> 00:10:19.430
If you open the files it makes in audacity,
you see what looks like a waveform of the sound.

00:10:19.430 --> 00:10:23.730
It is a waveform, but a waveform that’s
been plotted precisely on a graph.

00:10:23.730 --> 00:10:25.872
Zoom way, way, way in on the waveform,

00:10:25.872 --> 00:10:29.310
and eventually you can see the individual samples themselves.

00:10:29.310 --> 00:10:31.220
And that’s all digital sound is--

00:10:31.220 --> 00:10:33.960
it’s a huge list of numbers strung together in order.

00:10:33.960 --> 00:10:38.640
To get these numbers back into sound we can
hear, we need to use the opposite of an analog-to-digital

00:10:38.640 --> 00:10:40.020
converter, or “ADC”.

00:10:40.029 --> 00:10:43.470
So, we’ll use a DAC, or Digital-to-analog
converter.

00:10:43.470 --> 00:10:46.130
I like it when names make sense.

00:10:46.130 --> 00:10:50.180
A DAC will read the string of numbers, and
generate an analog voltage based upon their

00:10:50.180 --> 00:10:51.320
values.

00:10:51.320 --> 00:10:55.150
The DAC will smooth out the choppiness of
the samples a bit to make the resulting sound

00:10:55.150 --> 00:10:59.970
a little more natural, and now you’ve got
an analog signal to send into an amplifier

00:10:59.970 --> 00:11:01.920
and drive a loudspeaker.

00:11:01.920 --> 00:11:06.170
The result is a near-perfect reproduction
of the originally recorded sound.

00:11:06.170 --> 00:11:11.029
Here’s a very crude analogy to explain the
difference between analog and digital sound.

00:11:11.029 --> 00:11:16.000
A vinyl record’s walls generate an analog
signal by moving the stylus left and right...

00:11:16.000 --> 00:11:17.760
as well as up and down.

00:11:17.760 --> 00:11:21.250
It’s diagonally moved for stereo, but just
imagine for a moment that it’s just left

00:11:21.250 --> 00:11:22.250
and right.

00:11:22.250 --> 00:11:26.540
A record directly creates the analog signal
via the motion of the stylus.

00:11:26.540 --> 00:11:32.210
But a digital sound source is instead sort
of like a virtual stylus riding in a virtual groove.

00:11:32.210 --> 00:11:36.660
The sound samples are snapshots in time of
where the stylus was.

00:11:36.660 --> 00:11:41.820
A DAC will then create an analog signal by
running a virtual stylus through this virtual

00:11:41.820 --> 00:11:46.490
groove and placing it at exactly the correct
location--and thus generating the appropriate

00:11:46.490 --> 00:11:49.839
voltage level--as defined by the samples.

00:11:49.839 --> 00:11:54.040
By using a giant list of numbers to recreate
sound, instead of the physical properties

00:11:54.040 --> 00:11:59.480
of a plastic disc, the sound can be reproduced
flawlessly and accurately with no reliance

00:11:59.480 --> 00:12:03.670
on the record player’s cartridge properties,
the integrity of its stylus, it’s motor,

00:12:03.670 --> 00:12:05.800
the quality of the vinyl etc.

00:12:05.800 --> 00:12:10.220
The biggest boon of digital sound was that
it eliminated all of the little nuances that

00:12:10.220 --> 00:12:13.180
might change how a recording sounds.

00:12:13.180 --> 00:12:15.800
Digital sound is in a sense, absolute.

00:12:15.800 --> 00:12:20.730
But getting digital sound into the hands of
the average consumer took a long while.

00:12:20.730 --> 00:12:25.700
DACs and “ADCs” were expensive components,
and the amount of raw data generated by sound

00:12:25.700 --> 00:12:28.820
recording was immense for the standards of
the time.

00:12:28.820 --> 00:12:33.810
Although 650 megabytes, the data equivalent
of the first compact discs, is a paltry sum

00:12:33.810 --> 00:12:39.630
of data in the 21st century, it was unimaginably
huge in the early 1970’s, when the first

00:12:39.630 --> 00:12:42.370
commercial digital sound recording took place.

00:12:42.370 --> 00:12:48.270
For context, the Commodore 64, released the
same year as the compact disc, has 64 kilobytes

00:12:48.270 --> 00:12:51.850
of ram, and that was considered huge for the
time.

00:12:51.850 --> 00:12:55.930
A compact disc held roughly ten thousands
times as much data.

00:12:55.930 --> 00:13:00.080
64 kilobtyes of CD quality audio lasts this
long;

00:13:00.300 --> 00:13:00.800
(clip)

00:13:01.120 --> 00:13:02.620
That’s not super helpful.

00:13:02.620 --> 00:13:06.600
When we continue, we’ll look at the methods
that were used to store data from digital

00:13:06.600 --> 00:13:10.840
recordings, and we’ll discuss the rise of
the compact disc as a robust, consumer-friendly

00:13:10.840 --> 00:13:14.020
format for digital sound reproduction and
distribution.

00:13:14.020 --> 00:13:16.690
Thanks for watching, I hope you enjoyed the
video!

00:13:16.690 --> 00:13:20.170
If this is your first time coming across the
channel and you liked what you saw, please

00:13:20.170 --> 00:13:22.300
consider subscribing to Technology Connections.

00:13:22.300 --> 00:13:27.090
Don’t forget you can also follow me on Twitter
@TechConnectify, and you might enjoy the second

00:13:27.090 --> 00:13:31.790
channel, Technology Connection 2, where I
talk about stuff and don’t prepare for anything.

00:13:31.790 --> 00:13:37.279
Also, thanks to Lord Telaneo on Twitter, there
is also a Technology Connections Subreddit.

00:13:37.279 --> 00:13:42.770
I really don’t know reddit at all, but you
will also find me there as TechConnectify.

00:13:42.770 --> 00:13:46.870
As always, thank you to everyone who supports
this channel on Patreon, especially the wonderful

00:13:46.870 --> 00:13:49.320
folks that have been scrolling up your screen.

00:13:49.320 --> 00:13:52.700
It is with the support of people like you
that I’m able to make these videos.

00:13:52.700 --> 00:13:53.899
Thank you.

00:13:53.899 --> 00:13:57.060
If you’d like to you join these awesome
people and support the channel too, why not

00:13:57.060 --> 00:13:58.830
take a look at my Patreon page.

00:13:58.830 --> 00:14:01.600
Thank you for your consideration, and I’ll
see you next time!

